Papers by Per Erik Solberg
The Norwegian Dialect Corpus Treebank (2022.lrec-1)
Copied to clipboard
Andre Kåsen, Kristin Hagen, Anders Nøklestad, Joel Priestly, Per Erik Solberg, Dag Trygve Truslew Haug
| Challenge: | The NDC Treebank consists of recordings made between 2006 and 2012 and is annotated with morphological and syntactic information. |
| Approach: | They present the NDC Treebank of spoken Norwegian dialects in the Bokml variety of Norwegian. |
| Outcome: | The treebank consists of 4587 speech segments, overall 66009 tokens, from 17 different Norwegian dialects from south, west, east and north of Norway. |
The LIA Treebank of Spoken Norwegian Dialects (L18-1)
Copied to clipboard
Lilja Øvrelid, Andre Kåsen, Kristin Hagen, Anders Nøklestad, Per Erik Solberg, Janne Bondi Johannessen
| Challenge: | a long-term goal of this work is to develop a parser for spoken Norwegian with the immediate goal of parsing the whole LIA material. |
| Approach: | They describe the LIA treebank of transcribed spoken Norwegian dialects and their transcription, transliteration and further morphosyntactic annotation. |
| Outcome: | The treebank consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian. |
The Norwegian Parliamentary Speech Corpus (2022.lrec-1)
Copied to clipboard
| Challenge: | the dataset contains recordings of meetings at the Norwegian parliament . it is the first publicly available dataset containing unscripted, Norwegian speech . |
| Approach: | the Norwegian Parliamentary Speech Corpus is a publicly available speech dataset . it contains recordings of meetings from the Norwegian parliament with orthographic transcriptions . the dataset is intended to fill a gap in the available unscripted speech data . |
| Outcome: | the dataset contains recordings of meetings at the Norwegian parliament with orthographic transcriptions in Norwegian Bokml and Norwegian Nynorsk. |